Demand MemCpy: Overlapping of Computation and Data Transfer for Heterogeneous Computing
نویسندگان
چکیده
Heterogeneous computing relies on collaboration among different types of processors shared data. In systems with discrete accelerators (e.g., GP-GPU), data sharing requires transferring a large amount between CPU and accelerator memories can significantly increase the end-to-end execution time. This paper proposes novel mechanism called Demand MemCpy (DMC) to hide overheads. DMC copies from host memory based demands at page granularity. It utilizes hardware-only fetch requested short latency background pre-copy related pages in advance. Our evaluation shows that reduce time GP-GPU application by 25.4% average overlapping computation transfer not unused pages.
منابع مشابه
Computation of Seasonal Statistics from Annual Data for Iran's Economy
This article has no abstract.
متن کاملAsymptotic algorithm for computing the sample variance of interval data
The problem of the sample variance computation for epistemic inter-val-valued data is, in general, NP-hard. Therefore, known efficient algorithms for computing variance require strong restrictions on admissible intervals like the no-subset property or heavy limitations on the number of possible intersections between intervals. A new asymptotic algorithm for computing the upper bound of the samp...
متن کاملAutomatic Transformation for Overlapping Communication and Computation
Message-passing is a predominant programming paradigm for distributed memory systems. RDMA networks like infiniBand and Myrinet reduce communication overhead by overlapping communication with computation. For the overlap to be more effective, we propose a source-tosource transformation scheme by automatically restructuring message-passing codes. The extensions to control-flow graph can accurate...
متن کاملAccelerating Complex Data Transfer for Cluster Computing
The ability to move data quickly between the nodes of a distributed system is important for the performance of cluster computing frameworks, such as Hadoop and Spark. We show that in a cluster with modern networking technology data serialization is the main bottleneck and source of overhead in the transfer of rich data in systems based on high-level programming languages such as Java. We propos...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
ژورنال
عنوان ژورنال: IEEE Access
سال: 2022
ISSN: ['2169-3536']
DOI: https://doi.org/10.1109/access.2022.3195271